DeepGenResearch← All research
DeepGen Research · Paper · Model capability

Where language models go blind

A field guide to the limits of next-token prediction, and how to route around them.

Abstract

Large language models fail in structured, predictable ways wherever a task's native object is not a stream of language: exact computation, two-dimensional data, deep compositional logic, and self-checking. We argue these are not deficiencies to be prompted away but architectural mismatches, and that the reliable remedy is to route each operation to the mechanism whose inductive bias fits it — a specialized model, a deterministic tool, or a checked proposal — or to decline with a stated reason. We also separate two failures that are routinely conflated under the heading of “reasoning”, and show that they have opposite trajectories on current models and require different treatment.

1. One fact, and what follows from it

A large language model is trained to predict the next token in a sequence of language, and that training objective is the key to both its power and its limits. Where a task really is sequence-modelling over language — summarising, paraphrasing, drafting, translating — the model's prior is exactly the right one. Where the task's true object is something else, the same prior is not merely weaker; it is aimed at the wrong target.

This reframing matters because it changes what counts as a fix. If a model gets an arithmetic problem wrong, the instinct is to prompt it more carefully or to use a larger model. But arbitrary-precision arithmetic is not an emergent property of better next-token prediction; it is a different kind of computation. No prompt makes a language model into a calculator. The productive question is not “how do we make the model do this?” but “what mechanism should do this, and how does the model hand off to it?”

The blind spots are not bugs in a particular model. They are consequences of the objective every such model is trained on.

2. A map of the blind spots

Four regions recur across every current model. They differ in mechanism but share the same root: the native object is not a linguistic sequence.

Two-dimensional and spatial structure

A table is two-dimensional; the model reads a one-dimensional stream. Serialised into a line, a wide spreadsheet loses the correspondence between a value and its column, and the model's answers about it degrade as the table widens (Wu et al., 2025). The same blindness applies to fine spatial geometry: a model cannot triangulate over a dense set of raw coordinates it has been handed as text.

Example
“Across this 60-column export, list every account whose 2024 spend exceeds its 2023 spend, and give the regional totals.”

Exact numbers

Numbers are tokenised into fragments, and multi-digit arithmetic requires carries to propagate across those fragments in a way the architecture does not reliably support. Long decimals and scientific notation shred into meaningless pieces. The failure is quiet: the model returns a confident figure that is simply wrong.

Deep compositional logic

On problems whose difficulty grows by composition — more constraints, more steps, deeper nesting — accuracy does not decline gracefully. It holds, then collapses past a threshold (Shojaee et al., 2025). This is the region most often mistaken for a general verdict about “reasoning”, and §3 takes it apart.

Self-verification

A single forward pass emits tokens one at a time with no built-in check; the model cannot see its own error until the sequence is already written. Reasoning-style models and external verifiers add checks around a call, but a naive single call has none.

3. One failure that is really two

Careful sourcing matters most where a claim has travelled widely. In 2024, a study introduced controlled perturbations of grade-school maths problems and reported that adding a single irrelevant (“no-op”) clause could reduce a model's accuracy by as much as 65% — presented as evidence that models rely on pattern-matching rather than reasoning (Mirzadeh et al., 2025). The result was influential and is still cited as though it describes today's systems.

It largely does not. A 2026 replication ran the same test on current frontier models and found that the collapse reproduces only when the “irrelevant” clauses are left unaudited. Once they are filtered to genuinely irrelevant statements, the accuracy drop is statistically indistinguishable from zero (Sturgeon, 2026). The most plausible reading is that the models were making a reasonable inference — that information a questioner bothered to include might be relevant — rather than failing to reason.

Two failures, opposite trajectories

Distractor / surface-form fragility — adding an irrelevant clause, renaming a variable — was a real property of mid-2024 models and has largely resolved on the 2026 frontier. It needs no special handling beyond keeping a step focused on one mode of work.

Compositional-depth collapse — accuracy falling off a cliff as the logic deepens — is persistent, affects reasoning-optimised models too, and is not fixed by spending more inference-time compute; near the collapse point, models reduce their own effort (Shojaee et al., 2025). This one needs a solver.

The strongest published rebuttal to the collapse finding is, on inspection, an argument for exactly the fix we advocate. When models are asked to emit a compact procedure — a generating function — instead of an enumerated answer, they recover on instances previously scored as failures (Lawsen, 2025). That is the point: stop asking the model to grind out the answer in-context, have it produce the formal object, and let an engine execute it.

The lesson is not that models cannot reason. It is that thinking longer is not what fixes deep logic — offloading the search is.

4. The fix: route to the mechanism that fits

If the blind spots come from a mismatch between the task's object and the model's prior, the remedy is to detect the mismatch and route the operation to a mechanism whose bias is the right shape. Four destinations cover the cases.

Four destinations
  • Predict from context — where the answer can be inferred but not computed (a missing table value, a forecast), a specialized model built for that data type does the work.
  • Compute with a tool — where the answer is exact (arithmetic, aggregation, spatial queries, optimisation), a deterministic engine computes it and the model only sets up and explains.
  • Propose and check — where the model must generate something whose correctness it cannot self-verify (code, a schema-bound output, a proof), a checker validates before the result is trusted.
  • Decline with a reason — where there is no answerable ground truth, the system says so, rather than producing a confident guess.

Three of these keep the language model in the role it is actually good at — understanding the request and explaining the result — while the exact work happens elsewhere. The division of labour is the whole point. A specialized model earns its place only where the object is inferable-but-not-computable; wherever the object is exact, a tool or a solver holds it; and wherever the model produces something checkable, a checker confirms it. Often that checker is a piece of running code, not another model, because a decidable rule deserves a deterministic test.

5. Three worked examples

Exact figures. “Multiply these two 15-digit identifiers and apply a 4.5% adjustment.” The model writes the expression; an arbitrary-precision arithmetic tool computes it; the model reports the result. The carry problem disappears because the model never performs the arithmetic.

Scheduling under constraints. “Schedule 12 crews across 30 shifts under these 40 constraints, with no back-to-back nights and at least two crews per shift.” This is past the depth where a model reasons reliably, and more effort will not help. The model translates the problem into a formal constraint model; a solver searches and returns a certified schedule or proves none exists; the model explains it.

A wide, messy spreadsheet. “Reconcile these two 60-column sheets and flag rows where 2024 sales rose but expenses fell.” The grid is held in a store and queried, so the model never has to track columns in a flattened line. Where the layout itself carries meaning — merged headers on a scanned sheet — the structure is read visually and any numbers that reading produces are checked before they are used.

6. What remains hard

Honesty about the boundary is part of the method. Some things have no reliable mechanism today, and the right response is to say so rather than to route them somewhere that only appears to answer. Reading a complex visual table with a vision model, for instance, produces model-origin numbers that must still be checked; until a suitable checked path is in place, such outputs are treated as provisional. And a question with no fact of the matter — a genuinely subjective or normative one — is answered as a set of positions, not as a result. Naming these cases is what keeps the confident-but-wrong failure from reappearing at the edge.

A note on scope

This paper is the public version of an internal design document. The internal work uses a precise vocabulary for the four destinations and the checking rule, and grades every claim against its source. The aim here is the argument, not the notation; the claims that carry evidence are cited below.

References

  1. Mirzadeh, I. et al. (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. ICLR 2025. arXiv:2410.05229.
  2. Sturgeon, B. / LessWrong (2026). Revisiting GSM-Symbolic: Do 2026 Frontier Models Still Fail at Confounded Grade-School Math? — audited no-op distractor drop indistinguishable from zero on GPT-4o, Claude Opus 4.6, Claude Haiku 4.5.
  3. Shojaee, P. et al. (2025). The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.
  4. Lawsen, A. (2025). The Illusion of the Illusion of Thinking. arXiv:2506.09250.
  5. Wu, X., Ritter, A. & Xu, W. (2025). Tabular Data Understanding with Large Language Models: A Survey. arXiv:2508.00217.
  6. Anthropic (2026). Platform Documentation — token counting.